Goto

Collaborating Authors

 ieee asme transaction


Self-Closing Suction Grippers for Industrial Grasping via Form-Flexible Design

arXiv.org Artificial Intelligence

Shape-morphing robots have shown benefits in industrial grasping. We propose form-flexible grippers for adaptive grasping. The design is based on the hybrid jamming and suction mechanism, which deforms to handle objects that vary significantly in size from the aperture, including both larger and smaller parts. Compared with traditional grippers, the gripper achieves self-closing to form an airtight seal. Under a vacuum, a wide range of grasping is realized through the passive morphing mechanism at the interface that harmonizes pressure and flow rate. This hybrid gripper showcases the capability to securely grasp an egg, as small as 54.5% of its aperture, while achieving a maximum load-to-mass ratio of 94.3.


RationalVLA: A Rational Vision-Language-Action Model with Dual System

arXiv.org Artificial Intelligence

--A fundamental requirement for real-world robotic deployment is the ability to understand and respond to natural language instructions. Existing language-conditioned manipulation tasks typically assume that instructions are perfectly aligned with the environment. This assumption limits robustness and generalization in realistic scenarios where instructions may be ambiguous, irrelevant, or infeasible. T o address this problem, we introduce RAtional MAnipulation (RAMA), a new benchmark that challenges models with both unseen executable instructions and defective ones that should be rejected. In RAMA, we construct a dataset with over 14,000 samples, including diverse defective instructions spanning six dimensions: visual, physical, semantic, motion, safety, and out-of-context. We further propose the Rational Vision-Language-Action model (RationalVLA). It is a dual system for robotic arms that integrates the high-level vision-language model with the low-level manipulation policy by introducing learnable latent space embeddings. This design enables RationalVLA to reason over instructions, reject infeasible commands, and execute manipulation effectively. Experiments demonstrate that RationalVLA outperforms state-of-the-art baselines on RAMA by a 14.5% higher success rate and 0.94 average task length, while maintaining competitive performance on standard manipulation tasks. "Half of the troubles of this life can be traced to saying yes too quickly and not saying no soon enough. " -- Josh Billings Embodied intelligence represents the ultimate manifestation of artificial intelligence [1]. A necessary condition for the successful deployment of embodied intelligence in the real world is its ability to understand natural language and respond appropriately, either by providing correct answers or by executing the corresponding actions. This demand has sparked research on language-conditioned manipulation tasks, which require robots to follow natural language instructions to complete specific manipulation actions. Wenxuan Song, Jiayi Chen, Wenxue Li, Xu He, Xinhu Zheng, and Haoang Li are with The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China.


Robotic Grinding Skills Learning Based on Geodesic Length Dynamic Motion Primitives

arXiv.org Artificial Intelligence

--Learning grinding skills from human craftsmen by imitation learning has emerged as a prominent research topic in the field of robotic machining. Given their robust trajectory generalization ability and resilience to various external disturbances and environmental changes, Dynamical Movement Primitives (DMPs) provide a promising skills learning solution for the robotic grinding. However, challenges arise when directly applying DMPs to grinding tasks, including low orientation accuracy, inaccurate synchronization of position, orientation, and force, and the inability to generalize surface trajectories. T o address these issues, this paper proposes a robotic grinding skills learning method based on geodesic length DMPs (Geo-DMPs). First, a normalized two-dimensional weighted Gaussian kernel function and intrinsic mean clustering algorithm are proposed to extract surface geometric features from multiple demonstration trajectories. Then, an orientation manifold distance metric is introduced to exclude the time factor from the classical orientation DMPs, thereby constructing Geo-DMPs for the orientation learning to improve the orientation trajectory generation accuracy. On this basis, a synchronization encoding framework for position, orientation, and force skills is established, using a phase function related to geodesic length. This framework enables the generation of robotic grinding actions between any two points on the surface. Finally, experiments on robotic chamfer grinding and free-form surface grinding demonstrate that the proposed method exhibits high geometric accuracy and good generalization capabilities in encoding and generating grinding skills. This method holds significant implications for learning and promoting robotic grinding skills. T o the best of our knowledge, this may be the first attempt to use DMPs to generate grinding skills for position, orientation, and force on model-free surfaces, thereby presenting a novel approach to robotic grinding skills learning.


Generating Whole-Body Avoidance Motion through Localized Proximity Sensing

arXiv.org Artificial Intelligence

This paper presents a novel control algorithm for robotic manipulators in unstructured environments using proximity sensors partially distributed on the platform. The proposed approach exploits arrays of multi zone Time-of-Flight (ToF) sensors to generate a sparse point cloud representation of the robot surroundings. By employing computational geometry techniques, we fuse the knowledge of robot geometric model with ToFs sensory feedback to generate whole-body motion tasks, allowing to move both sensorized and non-sensorized links in response to unpredictable events such as human motion. In particular, the proposed algorithm computes the pair of closest points between the environment cloud and the robot links, generating a dynamic avoidance motion that is implemented as the highest priority task in a two-level hierarchical architecture. Such a design choice allows the robot to work safely alongside humans even without a complete sensorization over the whole surface. Experimental validation demonstrates the algorithm effectiveness both in static and dynamic scenarios, achieving comparable performances with respect to well established control techniques that aim to move the sensors mounting positions on the robot body. The presented algorithm exploits any arbitrary point on the robot surface to perform avoidance motion, showing improvements in the distance margin up to 100 mm, due to the rendering of virtual avoidance tasks on non-sensorized links.


Modular Adaptive Aerial Manipulation under Unknown Dynamic Coupling Forces

arXiv.org Artificial Intelligence

--Successful aerial manipulation largely depends on how effectively a controller can tackle the coupling dynamic forces between the aerial vehicle and the manipulator . However, this control problem has remained largely unsolved as the existing control approaches either require precise knowledge of the aerial vehicle/manipulator inertial couplings, or neglect the state-dependent uncertainties especially arising during the interaction phase. This work proposes an adaptive control solution to overcome this long standing control challenge without any a priori knowledge of the coupling dynamic terms. Additionally, in contrast to the existing adaptive control solutions, the proposed control framework is modular, that is, it allows independent tuning of the adaptive gains for the vehicle position sub-dynamics, the vehicle attitude sub-dynamics, and the manipulator sub-dynamics. Stability of the closed loop under the proposed scheme is derived analytically, and real-time experiments validate the effectiveness of the proposed scheme over the state-of-the-art approaches. I. INTRODUCTION An Unmanned Aerial Manipulator (UAM) is a coupled system where a quadrotor (or multirotor) vehicle carries a manipulator: the presence of the manipulator greatly improves the dexterity and flexibility of the quadrotor, making it capable to accomplish a wide range of tasks, from simple payload transportation to more complex tasks such as pick and place, contact-based inspection, grasping and assembling etc. [1]-[8]. This work was supported in part by "Aerial Manipulation" under IHFC grand project (GP/2021/DA/032), in part by "Capacity building for human resource development in Unmanned Aircraft System (Drone and related Technology)", MeiTY, India, in part by the Natural Science Foundation of China grants 62233004 and 62073074, and in part by Jiangsu Provincial Scientific Research Center of Applied Mathematics grant BK20233002.


Adaptive Visual Servoing for On-Orbit Servicing

arXiv.org Artificial Intelligence

This paper presents an adaptive visual servoing framework for robotic on-orbit servicing (OOS), specifically designed for capturing tumbling satellites. The vision-guided robotic system is capable of selecting optimal control actions in the event of partial or complete vision system failure, particularly in the short term. The autonomous system accounts for physical and operational constraints, executing visual servoing tasks to minimize a cost function. A hierarchical control architecture is developed, integrating a variant of the Iterative Closest Point (ICP) algorithm for image registration, a constrained noise-adaptive Kalman filter, fault detection and recovery logic, and a constrained optimal path planner. The dynamic estimator provides real-time estimates of unknown states and uncertain parameters essential for motion prediction, while ensuring consistency through a set of inequality constraints. It also adjusts the Kalman filter parameters adaptively in response to unexpected vision errors. In the event of vision system faults, a recovery strategy is activated, guided by fault detection logic that monitors the visual feedback via the metric fit error of image registration. The estimated/predicted pose and parameters are subsequently fed into an optimal path planner, which directs the robot's end-effector to the target's grasping point. This process is subject to multiple constraints, including acceleration limits, smooth capture, and line-of-sight maintenance with the target. Experimental results demonstrate that the proposed visual servoing system successfully captured a free-floating object, despite complete occlusion of the vision system.


Heuristic Predictive Control for Multi-Robot Flocking in Congested Environments

arXiv.org Artificial Intelligence

Multi-robot flocking possesses extraordinary advantages over a single-robot system in diverse domains, but it is challenging to ensure safe and optimal performance in congested environments. Hence, this paper is focused on the investigation of distributed optimal flocking control for multiple robots in crowded environments. A heuristic predictive control solution is proposed based on a Gibbs Random Field (GRF), in which bio-inspired potential functions are used to characterize robot-robot and robot-environment interactions. The optimal solution is obtained by maximizing a posteriori joint distribution of the GRF in a certain future time instant. A gradient-based heuristic solution is developed, which could significantly speed up the computation of the optimal control. Mathematical analysis is also conducted to show the validity of the heuristic solution. Multiple collision risk levels are designed to improve the collision avoidance performance of robots in dynamic environments. The proposed heuristic predictive control is evaluated comprehensively from multiple perspectives based on different metrics in a challenging simulation environment. The competence of the proposed algorithm is validated via the comparison with the non-heuristic predictive control and two existing popular flocking control methods. Real-life experiments are also performed using four quadrotor UAVs to further demonstrate the efficiency of the proposed design.


Design and Control of a Compact Series Elastic Actuator Module for Robots in MRI Scanners

arXiv.org Artificial Intelligence

In this study, we introduce a novel MRI-compatible rotary series elastic actuator module utilizing velocity-sourced ultrasonic motors for force-controlled robots operating within MRI scanners. Unlike previous MRI-compatible SEA designs, our module incorporates a transmission force sensing series elastic actuator structure, with four off-the-shelf compression springs strategically placed between the gearbox housing and the motor housing. This design features a compact size, thus expanding possibilities for a wider range of MRI robotic applications. To achieve precise torque control, we develop a controller that incorporates a disturbance observer tailored for velocity-sourced motors. This controller enhances the robustness of torque control in our actuator module, even in the presence of varying external impedance, thereby augmenting its suitability for MRI-guided medical interventions. Experimental validation demonstrates the actuator's torque control performance in both 3 Tesla MRI and non-MRI environments, achieving a settling time of 0.1 seconds and a steady-state error within 2% of its maximum output torque. Notably, our force controller exhibits consistent performance across low and high external impedance scenarios, in contrast to conventional controllers for velocity-sourced series elastic actuators, which struggle with steady-state performance under low external impedance conditions.


The Development of LLMs for Embodied Navigation

arXiv.org Artificial Intelligence

In recent years, the rapid advancement of Large Language Models (LLMs) such as the Generative Pre-trained Transformer (GPT) has attracted increasing attention due to their potential in a variety of practical applications. The application of LLMs with Embodied Intelligence has emerged as a significant area of focus. Among the myriad applications of LLMs, navigation tasks are particularly noteworthy because they demand a deep understanding of the environment and quick, accurate decision-making. LLMs can augment embodied intelligence systems with sophisticated environmental perception and decision-making support, leveraging their robust language and image-processing capabilities. This article offers an exhaustive summary of the symbiosis between LLMs and embodied intelligence with a focus on navigation. It reviews state-of-the-art models, research methodologies, and assesses the advantages and disadvantages of existing embodied navigation models and datasets. Finally, the article elucidates the role of LLMs in embodied intelligence, based on current research, and forecasts future directions in the field. A comprehensive list of studies in this survey is available at https://github.com/Rongtao-Xu/Awesome-LLM-EN


Adaptive Visual Servo Control for Autonomous Robots

arXiv.org Artificial Intelligence

This paper focuses on an adaptive and fault-tolerant vision-guided robotic system that enables to choose the most appropriate control action if partial or complete failure of the vision system in the short term occurs. Moreover, the autonomous robotic system takes physical and operational constraints into account to perform the demands of a specific visual servoing task in a way to minimize a cost function. A hierarchical control architecture is developed based on interwoven integration of a variant of the iterative closest point (ICP) image registration, a constrained noise-adaptive Kalman filter, a fault detection logic and recovery, together with a constrained optimal path planner. The dynamic estimator estimates unknown states and uncertain parameters required for motion prediction while imposing a set of inequality constraints for consistency of the estimation process and adjusting adaptively the Kalman filter parameters in the face of unexpected vision errors. It is followed by the implementation of a fault recovery strategy based on a fault detection logic that monitors the health of the visual feedback using the metric fit error of the image registration. Subsequently, the estimated/predicted pose and parameters are passed to an optimal path planner in order to bring the robot end-effector to the grasping point of a moving target as quickly as possible subject to multiple constraints such as acceleration limit, smooth capture, and line-of-sight angle of the target.